Skip to content

fix(bin): stop this home's own --resolve-key closes from re-waking the supervisor - #5008

Closed
tiago-peixoto wants to merge 10 commits into
kunchenguid:mainfrom
tiago-peixoto:fm/fm4907-sync
Closed

tiago-peixoto wants to merge 10 commits into
kunchenguid:mainfrom
tiago-peixoto:fm/fm4907-sync

Conversation

@tiago-peixoto

Copy link
Copy Markdown
Contributor

Intent

Clear base conflicts on #4907.
Branch fm/fm4885, last known head 0a21cd0.
The change is: stop this home's own --resolve-key answers from each waking the supervisor.
Was checks-green before conflicts. Do not open a second PR. Never merge upstream.

What Changed

  • Added a per-task home-owned append ledger to bin/fm-classify-lib.sh (state/.<task>.home-appends, v1 + identity header, merged half-open byte ranges, serialized by a sibling .lock), with status_home_appends_record / _covers / _ranges helpers; fm_wake_status_append_self_announced now records the exact range it appended before touching the watcher marker, and status_retire_presentation_task retires the ledger and its lock at teardown.
  • Split the seen predicate in bin/fm-wake-lib.sh: the marker-only check is now fm_wake_signal_reported_current, and fm_wake_signal_seen_current additionally treats growth past the classified offset as seen when every grown byte is in the owned ledger - so separate --resolve-key answers no longer each force a wake, while any foreign or unrecorded line still wakes. fm_wake_print_annotations keeps using the reported-only predicate, so owned closes still print in the signal annotation and UNREAD STATUS.
  • Covered the new behavior in tests/fm-wake-queue.test.sh, tests/fm-wake-drain-unread-status.test.sh, tests/fm-watch-triage.test.sh, and tests/fm-send-resolve-key.test.sh (repeat answers stay quiet across a regressed classified offset, a folded worker decision with no home append still wakes, owned growth still annotates turn-ended, unreadable logs still read as unreported), and documented the ledger in AGENTS.md, docs/architecture.md, and docs/scripts.md.

Risk Assessment

✅ Low: The change is well-bounded to one new per-task sidecar plus a split wake/presentation predicate, every path I traced fails toward waking (worker bytes always land in an uncovered gap, and identity mismatch, missing ledger, unreadable file, and shrunk file all read as unreported), the captain-facing presentation path is restored byte-identically to base, retirement is complete, the discriminating regression tests fail on base, and the branch still merges cleanly with current origin/main.

Testing

I stood up throwaway firstmate homes in /tmp and drove the supervision loop the way an operator does: a worker opens keyed decisions, the real watcher process wakes and exits, the operator drains and acknowledges, then answers with real fm-send --resolve-key calls. The decisive evidence is a live before/after of the actual failure. The watcher captures a status file's classified endpoint at the top of a poll and only commits it after a slow crew-evidence subprocess; an answer written inside that window leaves the seen marker not vouching for the answer's own bytes. Forcing that window (a deliberately slow crew-state probe, which is what the real probe is), base 1b1b6e0 exits with signal: .../t1.status - the supervisor woken by its own close - and target a3624c9 stays asleep, then still wakes on the worker's next blocked: line. The same run proves the presentation guard the author was asked for in review round 2: the turn-ended wake's historical annotation still prints t1.status: resolved [key=budget]: answered: approved, go ahead. A second driver covers the ordinary path end to end and finishes with a real bin/fm-teardown.sh that leaves no home-appends ledger or lock behind. A third hammers the new lock with 12 simultaneous real fm-send --resolve-key processes: the ledger stays a single well-formed ascending range inside the log's bounds, every landed close is presented, and nothing is silently lost. While writing that one I hit sends failing with "task metadata could not be locked for final delivery validation"; that reproduces identically on base under the same burst, is fm-send's own fail-closed meta lock rather than anything this change touches, leaves the decision open and tells the operator, so I scoped the scenario to what the change owns instead of reporting it against this PR. This product has no graphical surface - the end-user surfaces are the watcher's reason lines and the drain's rendered sections - so the artifacts are CLI transcripts of those exact surfaces rather than screenshots. Of the four suites the change touches, three ran green in full; tests/fm-watch-triage.test.sh was still crawling at 98 passes with zero failures when I finished, starved by a parallel run of the same suite in another no-mistakes worktree on this host, and every case this change adds or touches in it had already passed. The before/after of the new regression test itself was exercised only as a test-suite run, not against the running product, so it is reported as untested here; the live before/after of the same defect through the real watcher covers that ground. Broad regression across the rest stays with CI.

  • Live validation: ✅ go - 7 of 9 scenarios driven live against the product
Scenario Result Live Evidence
An in-flight watcher classification lands after the supervisor's --resolve-key answer; the supervisor must not be woken by its own close ✅ pass live live-inflight-classify-race.sh run against both commits (race-before-after.txt): base exits signal: .../t1.status, target stays asleep. The driver first asserts the race actually reproduced (class…
Two separate --resolve-key answers on one task wake the supervisor zero extra times ✅ pass live live-resolve-key-wake.sh scenario 4: after two real fm-send --resolve-key answers, the real fm-watch.sh survives a full poll cycle with no wake reason printed and an empty durable .wake-queue (live-…
Adversarial: a worker-authored line after the owned answers still wakes the supervisor ✅ pass live live-resolve-key-wake.sh scenario 5 and live-inflight-classify-race.sh step 8: appending blocked: need staging credentials makes the real watcher exit with signal: .../t1.status
Guard: the turn-ended historical annotation still presents this home's own close (recorded decision review-r2-4) ✅ pass live live-inflight-classify-race.sh step 7 on target: with the wake suppressed, the turn-end marker wakes the watcher and the drain prints `wake annotation: ... t1.status: resolved [key=budget]: answered…
Guard: UNREAD STATUS and the signal annotation still print both owned closes ✅ pass live live-resolve-key-wake.sh scenario 6: the drain prints both resolved [key=budget] and resolved [key=vendor] annotations alongside the worker's blocker
Adversarial: a real fm-teardown.sh leaves no orphaned home-appends ledger or lock ✅ pass live live-resolve-key-wake.sh scenario 7: a real project clone, worktree and task are torn down with bin/fm-teardown.sh task-x1 (exit 0) after a real --resolve-key answer wrote the ledger and a stale `…
Adversarial: simultaneous --resolve-key answers on one task keep the ledger well formed and lose no close ✅ pass live live-concurrent-answers.sh with 12 concurrent real fm-send processes, 5 consecutive target runs: single merged ascending range inside the log's byte bounds, no torn line, every landed close presente…
The change's regression test reproduces the reported failure: fails before the fix, passes after ⏸️ untested no The prior payload recorded this scenario as live=false: it rests only on running the unit test file tests/fm-send-resolve-key.test.sh (copied onto base 1b1b6e0 and on target), which is a test harnes…
Full tests/fm-watch-triage.test.sh regression pass ⏸️ untested no The suite is sleep-heavy and was starved by a parallel run of the same file in another no-mistakes worktree on this host, so it could not finish inside this step (98 of ~126 cases had passed with zero…
Evidence: Live before/after: the supervisor woken by its own --resolve-key answer on base, silent on target

Source: Live before/after: the supervisor woken by its own --resolve-key answer on base, silent on target

########## BASE 1b1b6e0 (before the fix) ########## PASS watcher's classified offset regressed to 92, behind the answer's bytes - the race reproduced FAIL the supervisor was WOKEN by its own --resolve-key answer: signal: .../home/state/t1.status base EXIT=1 ########## TARGET a3624c9 (with the fix) ########## PASS watcher's classified offset regressed to 92, behind the answer's bytes - the race reproduced PASS the watcher stayed asleep: this home's own answer did not wake it PASS the worker's turn end woke the supervisor: signal: .../home/state/t1.turn-ended PASS the turn-ended annotation still presents this home's own close PASS the worker's blocked: line woke the supervisor: signal: .../home/state/t1.status target EXIT=0

########## BASE 1b1b6e0 (before the fix) ##########
checkout under test: /tmp/fm-base-1b1b6e0.hrZQQi
throwaway FM_HOME:   /tmp/fm-live-race.sSvxaW/home
crew-probe hold:     8s

=== step 1: the worker's decision wakes the supervisor (the legitimate wake)
PASS  woken: signal: /tmp/fm-live-race.sSvxaW/home/state/t1.status
PASS  wake acknowledged

=== step 2: the worker keeps working (routine growth the watcher will classify)
status size before the answer: 92 bytes

=== step 3: the watcher enters its slow crew-evidence check
PASS  watcher is inside the crew-state probe with its endpoint already captured

=== step 4: the supervisor answers the decision INSIDE that window
PASS  answer sent while the watcher was mid-check
status size after the answer: 144 bytes

=== step 5: the watcher commits the endpoint it captured before the answer
PASS  watcher's classified offset regressed to 92, behind the answer's bytes - the race reproduced

=== step 6 (the question): is the supervisor woken by its own answer?
FAIL  the supervisor was WOKEN by its own --resolve-key answer: signal: /tmp/fm-live-race.sSvxaW/home/state/t1.status

=== step 7 (guard): the worker's turn ends - the owned close must still be annotated
FAIL  the watcher was already gone before the turn-end guard

=== step 8 (adversarial): a real worker line must still wake
FAIL  a real worker line after the owned answer was swallowed: check: rearm-resurface

=== step 9 (guard): the answer is still on the captain-facing surface
1789878685	4	signal	t1.status	signal: /tmp/fm-live-race.sSvxaW/home/state/t1.status
wake annotation: unread wake-EVENT since last drain, not current state: t1.status: working: still refactoring the adapter
wake annotation: unread wake-EVENT since last drain, not current state: t1.status: resolved [key=budget]: answered: approved, go ahead
wake annotation: latest wake-EVENT observed at drain, not current state: t1.status: blocked: need staging credentials to continue
OPEN DECISIONS (still open, folded from the durable status logs - not just the latest line):
t1 blocked: need staging credentials to continue
OPEN DECISIONS: close one by answering it: bin/fm-send.sh <task> --resolve-key <key> '<answer>'
PASS  the worker's blocker is presented to the captain

===========================
RESULT: at least one live scenario FAILED
base EXIT=1

########## TARGET a3624c9 (with the fix) ##########
checkout under test: ~/.no-mistakes/worktrees/42fe67ed39fa/01M2YF86AG8VETDN9HQKHA0P9N
throwaway FM_HOME:   /tmp/fm-live-race.CVBoZv/home
crew-probe hold:     8s

=== step 1: the worker's decision wakes the supervisor (the legitimate wake)
PASS  woken: signal: /tmp/fm-live-race.CVBoZv/home/state/t1.status
PASS  wake acknowledged

=== step 2: the worker keeps working (routine growth the watcher will classify)
status size before the answer: 92 bytes

=== step 3: the watcher enters its slow crew-evidence check
PASS  watcher is inside the crew-state probe with its endpoint already captured

=== step 4: the supervisor answers the decision INSIDE that window
PASS  answer sent while the watcher was mid-check
status size after the answer: 144 bytes

=== step 5: the watcher commits the endpoint it captured before the answer
PASS  watcher's classified offset regressed to 92, behind the answer's bytes - the race reproduced

=== step 6 (the question): is the supervisor woken by its own answer?
PASS  the watcher stayed asleep: this home's own answer did not wake it

=== step 7 (guard): the worker's turn ends - the owned close must still be annotated
PASS  the worker's turn end woke the supervisor: signal: /tmp/fm-live-race.CVBoZv/home/state/t1.turn-ended
1789878738	4	signal	t1.turn-ended	signal: /tmp/fm-live-race.CVBoZv/home/state/t1.turn-ended
wake annotation: unread wake-EVENT since last drain, not current state; historical / not necessarily the triggering event: t1.status: working: still refactoring the adapter
wake annotation: latest wake-EVENT observed at drain, not current state; historical / not necessarily the triggering event: t1.status: resolved [key=budget]: answered: approved, go ahead
PASS  the turn-ended annotation still presents this home's own close

=== step 8 (adversarial): a real worker line must still wake
PASS  the worker's blocked: line woke the supervisor: signal: /tmp/fm-live-race.CVBoZv/home/state/t1.status

=== step 9 (guard): the answer is still on the captain-facing surface
1789878749	6	signal	t1.status	signal: /tmp/fm-live-race.CVBoZv/home/state/t1.status
wake annotation: unread wake-EVENT since last drain, not current state: t1.status: working: still refactoring the adapter
wake annotation: unread wake-EVENT since last drain, not current state: t1.status: resolved [key=budget]: answered: approved, go ahead
wake annotation: latest wake-EVENT observed at drain, not current state: t1.status: blocked: need staging credentials to continue
OPEN DECISIONS (still open, folded from the durable status logs - not just the latest line):
t1 blocked: need staging credentials to continue
OPEN DECISIONS: close one by answering it: bin/fm-send.sh <task> --resolve-key <key> '<answer>'
PASS  the worker's blocker is presented to the captain

===========================
RESULT: all live scenarios passed
target EXIT=0
Evidence: Driver: live in-flight-classification race through the real watcher, fm-send and drain

Source: Driver: live in-flight-classification race through the real watcher, fm-send and drain

#!/usr/bin/env bash
# Live reproduction of the failure PR #4907 fixes, driven through the REAL
# watcher process (bin/fm-watch.sh) and the REAL answer command (bin/fm-send.sh).
#
# The race, which happens in ordinary operation: the watcher captures a status
# file's classified endpoint at the top of a poll, then spends real time on its
# "is the crew provably working?" evidence check (a bounded no-mistakes call).
# If the supervisor answers a decision with --resolve-key inside that window,
# the watcher afterwards commits the endpoint it captured BEFORE the answer. The
# watcher's seen marker no longer vouches for the answer's own bytes, so on the
# next poll the supervisor is woken - by nothing but its own bookkeeping close.
#
# This driver forces that window deterministically by making the crew-state
# probe slow, which is exactly what the real probe is (a subprocess call), and
# then asks the only question that matters to a user: does the watcher exit?
set -u

ROOT=${FM_LIVE_ROOT:?set FM_LIVE_ROOT to the checkout under test}
HOLD=${FM_RACE_HOLD:-8}
WORK=$(mktemp -d "${TMPDIR:-/tmp}/fm-live-race.XXXXXX")
HOMEDIR="$WORK/home"; STATE="$HOMEDIR/state"; FAKEBIN="$WORK/fakebin"
mkdir -p "$STATE" "$FAKEBIN" "$WORK/notangle"
STATUS="$STATE/t1.status"
MODE="$WORK/crew-mode"; MARKER="$WORK/crew-held"
FAILED=0
say() { printf '\n=== %s\n' "$*"; }
ok()  { printf 'PASS  %s\n' "$*"; }
bad() { printf 'FAIL  %s\n' "$*"; FAILED=1; }

cat > "$FAKEBIN/tmux" <<'SH'
#!/usr/bin/env bash
set -u
case "${1:-}" in
  send-keys)
    shift; literal=0
    while [ $# -gt 0 ]; do
      case "$1" in
        -t) shift 2 ;;
        -l) literal=1; shift ;;
        *) break ;;
      esac
    done
    [ "$literal" = 1 ] && printf '%s' "${1:-}" >> "${FM_SEND_LOG:-/dev/null}"
    exit 0 ;;
  display-message)
    for a in "$@"; do case "$a" in *cursor_y*) printf '1\n'; exit 0 ;; esac; done
    printf 'fakepane\n'; exit 0 ;;
  capture-pane)
    # A live pane: its content changes on every capture, so the unrelated
    # stale-pane supervision layer never fires and this driver observes only
    # the status SIGNAL path, which is the layer the ledger touches.
    n=$(cat "${FM_RACE_PANE_COUNTER:-/dev/null}" 2>/dev/null || printf 0)
    [ -n "${FM_RACE_PANE_COUNTER:-}" ] && printf '%s\n' "$((n + 1))" > "$FM_RACE_PANE_COUNTER"
    printf '╭────╮\n│ %s │\n╰────╯\n' "$((n + 1))"
    exit 0 ;;
  list-windows)
    printf '%s\n' fm-t1; exit 0 ;;
esac
exit 0
SH
# The crew-state probe: a real subprocess, deliberately slow in "hold" mode so
# the watcher's evidence check spans the moment the answer is written.
cat > "$FAKEBIN/fm-crew-state.sh" <<'SH'
#!/usr/bin/env bash
mode=$(cat "$FM_RACE_MODE" 2>/dev/null || printf idle)
case "$mode" in
  hold)
    printf 'held\n' >> "$FM_RACE_MARKER"
    /bin/sleep "${FM_RACE_HOLD:-8}"
    printf 'state: working · source: run-step · mid no-mistakes step\n' ;;
  working) printf 'state: working · source: run-step · mid no-mistakes step\n' ;;
  *) printf 'state: unknown · source: none · idle worker\n' ;;
esac
exit 0
SH
printf '#!/usr/bin/env bash\nexit 0\n' > "$WORK/wedge-rec"
chmod +x "$FAKEBIN/tmux" "$FAKEBIN/fm-crew-state.sh" "$WORK/wedge-rec"
export FM_WEDGE_ALARM_EXEC="$WORK/wedge-rec"
export FM_ROOT_OVERRIDE="$WORK/notangle"
export FM_RACE_MODE="$MODE" FM_RACE_MARKER="$MARKER" FM_RACE_HOLD="$HOLD"

DRAIN="$ROOT/bin/fm-wake-drain.sh"; WATCH="$ROOT/bin/fm-watch.sh"; SEND="$ROOT/bin/fm-send.sh"
size_of() { LC_ALL=C wc -c < "$1" | tr -d '[:space:]'; }
classified_offset() {
  FM_STATE_OVERRIDE="$STATE" bash -c '. "$1"; fm_wake_signal_seen_size "$2" "$3"' \
    _ "$ROOT/bin/fm-wake-lib.sh" "$STATE" "$STATUS"
}
ack_cycle() {
  local err seq gen; err="$WORK/ack.err"
  FM_STATE_OVERRIDE="$STATE" "$DRAIN" >/dev/null 2>"$err" || return 1
  seq=$(sed -n 's/^WAKE_ACK_REQUIRED:.*--ack-through \([0-9][0-9]*\) --recovery-generation [A-Za-z0-9._-][A-Za-z0-9._-]*$/\1/p' "$err")
  gen=$(sed -n 's/^WAKE_ACK_REQUIRED:.*--ack-through [0-9][0-9]* --recovery-generation \([A-Za-z0-9._-][A-Za-z0-9._-]*\)$/\1/p' "$err")
  [ -n "$seq" ] && [ -n "$gen" ] || return 1
  FM_STATE_OVERRIDE="$STATE" "$DRAIN" --ack-through "$seq" --recovery-generation "$gen"
}
watch_bg() {
  PATH="$FAKEBIN:$PATH" FM_STATE_OVERRIDE="$STATE" FM_CREW_STATE_BIN="$FAKEBIN/fm-crew-state.sh" \
    FM_RACE_PANE_COUNTER="$WORK/pane-counter" \
    FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \
    "$WATCH" > "$1" 2>"$1.err" &
}
wait_for_exit() { local pid=$1 limit=$2 i=0
  while [ "$i" -lt "$limit" ]; do kill -0 "$pid" 2>/dev/null || return 0; sleep 0.1; i=$((i+1)); done
  return 1; }
reap() { kill "$1" 2>/dev/null || true; wait "$1" 2>/dev/null || true; }
send() {
  env -u NO_MISTAKES_GATE PATH="$FAKEBIN:$PATH" FM_GATE_REFUSE_BYPASS=1 \
    FM_ROOT_OVERRIDE="$HOMEDIR" FM_HOME="$HOMEDIR" FM_SEND_LOG="$WORK/send.log" FM_SEND_SETTLE=0 \
    "$SEND" t1 --resolve-key "$1" "$2"
}

printf 'checkout under test: %s\n' "$ROOT"
printf 'throwaway FM_HOME:   %s\n' "$HOMEDIR"
printf 'crew-probe hold:     %ss\n' "$HOLD"

printf idle > "$MODE"
printf 'window=sess:fm-t1\nkind=ship\n' > "$STATE/t1.meta"
printf 'needs-decision [key=budget]: approve the $400 spend?\n' > "$STATUS"

say "step 1: the worker's decision wakes the supervisor (the legitimate wake)"
watch_bg "$WORK/w1.out"; W1=$!
wait_for_exit "$W1" 150 && grep -qF "signal: $STATUS" "$WORK/w1.out" \
  && ok "woken: $(cat "$WORK/w1.out")" || bad "the decision never woke the supervisor"
reap "$W1"
ack_cycle && ok "wake acknowledged" || bad "could not acknowledge the wake"

say "step 2: the worker keeps working (routine growth the watcher will classify)"
printf 'working: still refactoring the adapter\n' >> "$STATUS"
SIZE2=$(size_of "$STATUS")
printf 'status size before the answer: %s bytes\n' "$SIZE2"

say "step 3: the watcher enters its slow crew-evidence check"
printf hold > "$MODE"
watch_bg "$WORK/w2.out"; W2=$!
i=0; while [ "$i" -lt 300 ] && [ ! -s "$MARKER" ]; do sleep 0.1; i=$((i+1)); done
[ -s "$MARKER" ] && ok "watcher is inside the crew-state probe with its endpoint already captured" \
  || bad "the watcher never reached the crew-evidence check"

say "step 4: the supervisor answers the decision INSIDE that window"
send budget "approved, go ahead" && ok "answer sent while the watcher was mid-check" || bad "the --resolve-key send failed"
printf 'status size after the answer: %s bytes\n' "$(size_of "$STATUS")"

say "step 5: the watcher commits the endpoint it captured before the answer"
printf working > "$MODE"
i=0
while [ "$i" -lt $(( (HOLD + 12) * 10 )) ]; do
  [ "$(classified_offset)" = "$SIZE2" ] && break
  kill -0 "$W2" 2>/dev/null || break
  sleep 0.1; i=$((i+1))
done
if [ "$(classified_offset)" = "$SIZE2" ]; then
  ok "watcher's classified offset regressed to $SIZE2, behind the answer's bytes - the race reproduced"
else
  bad "the race did not reproduce (classified offset is $(classified_offset), wanted $SIZE2); raise FM_RACE_HOLD and retry"
fi

say "step 6 (the question): is the supervisor woken by its own answer?"
printf idle > "$MODE"
if wait_for_exit "$W2" 200; then
  bad "the supervisor was WOKEN by its own --resolve-key answer: $(cat "$WORK/w2.out")"
else
  ok "the watcher stayed asleep: this home's own answer did not wake it"
fi

say "step 7 (guard): the worker's turn ends - the owned close must still be annotated"
# The wake is suppressed, but presentation must not be. A turn-ended wake row is
# the captain-facing surface the owned-append ledger is forbidden to touch.
if kill -0 "$W2" 2>/dev/null; then
  : > "$STATE/t1.turn-ended"
  if wait_for_exit "$W2" 200; then
    ok "the worker's turn end woke the supervisor: $(grep -F 'signal:' "$WORK/w2.out" | tail -n 1)"
  else
    bad "the turn-end marker did not wake the supervisor"
  fi
  reap "$W2"
  FM_STATE_OVERRIDE="$STATE" "$DRAIN" > "$WORK/drain-turnend.out" 2>"$WORK/drain-turnend.err" || true
  cat "$WORK/drain-turnend.out"
  grep -qF 'resolved [key=budget]: answered: approved, go ahead' "$WORK/drain-turnend.out" \
    && ok "the turn-ended annotation still presents this home's own close" \
    || bad "the ledger hid the owned close from the turn-ended annotation"
  ack_cycle >/dev/null 2>&1 || true
else
  bad "the watcher was already gone before the turn-end guard"
fi

say "step 8 (adversarial): a real worker line must still wake"
printf 'blocked: need staging credentials to continue\n' >> "$STATUS"
watch_bg "$WORK/w3.out"; W3=$!
wait_for_exit "$W3" 250 && grep -qF "signal: $STATUS" "$WORK/w3.out" \
  && ok "the worker's blocked: line woke the supervisor: $(grep -F 'signal:' "$WORK/w3.out" | tail -n 1)" \
  || bad "a real worker line after the owned answer was swallowed: $(cat "$WORK/w3.out")"
reap "$W3" 2>/dev/null || true

say "step 9 (guard): the answer is still on the captain-facing surface"
FM_STATE_OVERRIDE="$STATE" "$DRAIN" > "$WORK/drain.out" 2>"$WORK/drain.err" || true
cat "$WORK/drain.out"
grep -qF 'blocked: need staging credentials' "$WORK/drain.out" \
  && ok "the worker's blocker is presented to the captain" \
  || bad "the worker's blocker was not presented"

printf '\n===========================\n'
[ "$FAILED" -eq 0 ] && printf 'RESULT: all live scenarios passed\n' || printf 'RESULT: at least one live scenario FAILED\n'
exit "$FAILED"
Evidence: Live supervisor loop on target: two answers, no re-wake, presentation intact, real fm-teardown.sh retires the ledger

Source: Live supervisor loop on target: two answers, no re-wake, presentation intact, real fm-teardown.sh retires the ledger

=== scenario 4 (the fix): neither of this home's own answers wakes the supervisor again PASS watcher stayed asleep through a full poll cycle: no wake reason, empty wake queue === scenario 5 (adversarial): a later worker line on the same task still wakes PASS the worker's blocked: line woke the supervisor: signal: .../home/state/t1.status === scenario 6 (guard): the answers are still shown to the captain, not hidden wake annotation: unread wake-EVENT since last drain ...: t1.status: resolved [key=budget]: answered: approved, go ahead wake annotation: unread wake-EVENT since last drain ...: t1.status: resolved [key=vendor]: answered: go with vendor B PASS both owned closes still printed on the captain-facing surface === scenario 7 (adversarial): a real fm-teardown.sh leaves no orphaned ledger state PASS precondition: the answered task carries a live home-appends ledger fm-teardown.sh exit=0 PASS real fm-teardown.sh removed the per-task ledger and its stale lock remaining task-x1 state: none

checkout under test: ~/.no-mistakes/worktrees/42fe67ed39fa/01M2YF86AG8VETDN9HQKHA0P9N (a3624c9)
throwaway FM_HOME:   /tmp/fm-live-4907.zPzfqz/home

=== setup: a ship worker opens two captain decisions on task t1
needs-decision [key=budget]: approve the $400 spend?
needs-decision [key=vendor]: vendor A or vendor B?

=== scenario 1: the worker's two decisions wake the supervisor once
PASS  watcher exited and woke the supervisor: signal: /tmp/fm-live-4907.zPzfqz/home/state/t1.status

=== scenario 2: the supervisor sees both decisions and acknowledges the wake
1789878172	2	signal	t1.status	needs-decision: /tmp/fm-live-4907.zPzfqz/home/state/t1.status
wake annotation: unread wake-EVENT since last drain, not current state: t1.status: needs-decision [key=budget]: approve the $400 spend?
wake annotation: latest wake-EVENT observed at drain, not current state: t1.status: needs-decision [key=vendor]: vendor A or vendor B?
OPEN DECISIONS (still open, folded from the durable status logs - not just the latest line):
t1 [key=budget] needs-decision: approve the $400 spend?
t1 [key=vendor] needs-decision: vendor A or vendor B?
OPEN DECISIONS: close one by answering it: bin/fm-send.sh <task> --resolve-key <key> '<answer>'
PASS  both open decisions presented to the captain
PASS  wake acknowledged; queue is empty

=== scenario 3: the supervisor answers BOTH decisions with two --resolve-key sends
PASS  first answer sent (key=budget)
PASS  second answer sent (key=vendor)
--- status log after both answers ---
needs-decision [key=budget]: approve the $400 spend?
needs-decision [key=vendor]: vendor A or vendor B?
resolved [key=budget]: answered: approved, go ahead
resolved [key=vendor]: answered: go with vendor B

=== scenario 4 (the fix): neither of this home's own answers wakes the supervisor again
PASS  watcher stayed asleep through a full poll cycle: no wake reason, empty wake queue

=== scenario 5 (adversarial): a later worker line on the same task still wakes
PASS  the worker's blocked: line woke the supervisor: signal: /tmp/fm-live-4907.zPzfqz/home/state/t1.status

=== scenario 6 (guard): the answers are still shown to the captain, not hidden
1789878183	4	signal	t1.status	signal: /tmp/fm-live-4907.zPzfqz/home/state/t1.status
wake annotation: unread wake-EVENT since last drain, not current state: t1.status: resolved [key=budget]: answered: approved, go ahead
wake annotation: unread wake-EVENT since last drain, not current state: t1.status: resolved [key=vendor]: answered: go with vendor B
wake annotation: latest wake-EVENT observed at drain, not current state: t1.status: blocked: need staging credentials to continue
OPEN DECISIONS (still open, folded from the durable status logs - not just the latest line):
t1 blocked: need staging credentials to continue
OPEN DECISIONS: close one by answering it: bin/fm-send.sh <task> --resolve-key <key> '<answer>'
PASS  both owned closes still printed on the captain-facing surface
PASS  the worker's blocker is presented too

=== scenario 7 (adversarial): a real fm-teardown.sh leaves no orphaned ledger state
PASS  precondition: the answered task carries a live home-appends ledger
.task-x1.home-appends
fm-teardown.sh exit=0
==> /tmp/fm-live-4907.zPzfqz/td.out <==
teardown task-x1 complete (window firstmate:fm-task-x1, worktree /tmp/fm-live-4907.zPzfqz/td/wt)
Backlog: task-x1 just finished (this home keeps no markdown backlog at /tmp/fm-live-4907.zPzfqz/td/data/backlog.md). Update /tmp/fm-live-4907.zPzfqz/td/data/backlog.md - move task-x1 to Done, keep Done to the 10 most recent, then re-scan Queued and dispatch only work whose blockers are gone and date is due.

==> /tmp/fm-live-4907.zPzfqz/td.err <==
PASS  real fm-teardown.sh removed the per-task ledger and its stale lock
remaining task-x1 state: none

===========================
RESULT: all live scenarios passed
artifacts under /tmp/fm-live-4907.zPzfqz
Evidence: Driver: live end-to-end supervisor wake loop including real teardown

Source: Driver: live end-to-end supervisor wake loop including real teardown

#!/usr/bin/env bash
# Live end-to-end driver for PR #4907's intent:
# "stop this home's own --resolve-key answers from each waking the supervisor".
#
# Drives the REAL product executables in a throwaway FM_HOME:
#   bin/fm-watch.sh       the watcher that wakes the supervisor (exits on an
#                         actionable wake, printing its reason)
#   bin/fm-send.sh        the supervisor answering a decision with --resolve-key
#   bin/fm-wake-drain.sh  the captain-facing presentation of that wake
#
# Nothing here reads the implementation source. Every assertion is on observable
# product behaviour: whether the watcher process exits (= the supervisor is
# woken), what the durable wake queue holds, and what the drain prints.
set -u

ROOT=${FM_LIVE_ROOT:?set FM_LIVE_ROOT to the checkout under test}
WORK=$(mktemp -d "${TMPDIR:-/tmp}/fm-live-4907.XXXXXX")
HOMEDIR="$WORK/home"; STATE="$HOMEDIR/state"; FAKEBIN="$WORK/fakebin"
mkdir -p "$STATE" "$FAKEBIN" "$WORK/notangle"
STATUS="$STATE/t1.status"

step=0
say() { printf '\n=== %s\n' "$*"; }
ok()  { printf 'PASS  %s\n' "$*"; }
bad() { printf 'FAIL  %s\n' "$*"; FAILED=1; }
FAILED=0

# --- stubs: the backend pane transport and the crew-state reader -------------
# These stand in for a real terminal multiplexer and a real no-mistakes run
# probe. Everything under test (watcher triage, ledger, drain) is the real code.
cat > "$FAKEBIN/tmux" <<'SH'
#!/usr/bin/env bash
set -u
case "${1:-}" in
  send-keys)
    shift; literal=0
    while [ $# -gt 0 ]; do
      case "$1" in
        -t) shift 2 ;;
        -l) literal=1; shift ;;
        *) break ;;
      esac
    done
    [ "$literal" = 1 ] && printf '%s' "${1:-}" >> "${FM_SEND_LOG:-/dev/null}"
    exit 0 ;;
  display-message)
    for a in "$@"; do case "$a" in *cursor_y*) printf '1\n'; exit 0 ;; esac; done
    printf 'fakepane\n'; exit 0 ;;
  capture-pane) printf '╭────╮\n│    │\n╰────╯\n'; exit 0 ;;
  list-windows) printf '%s\n' fm-t1; exit 0 ;;
esac
exit 0
SH
cat > "$FAKEBIN/fm-crew-state.sh" <<'SH'
#!/usr/bin/env bash
printf 'state: unknown · source: none · idle worker\n'
exit 0
SH
cat > "$FAKEBIN/sleep" <<'SH'
#!/usr/bin/env bash
exit 0
SH
cat > "$WORK/wedge-rec" <<'SH'
#!/usr/bin/env bash
exit 0
SH
chmod +x "$FAKEBIN/tmux" "$FAKEBIN/fm-crew-state.sh" "$FAKEBIN/sleep" "$WORK/wedge-rec"

export FM_WEDGE_ALARM_EXEC="$WORK/wedge-rec"
export FM_ROOT_OVERRIDE="$WORK/notangle"

DRAIN="$ROOT/bin/fm-wake-drain.sh"
WATCH="$ROOT/bin/fm-watch.sh"
SEND="$ROOT/bin/fm-send.sh"

drain() { FM_STATE_OVERRIDE="$STATE" "$DRAIN" "$@"; }

ack_cycle() {  # drain, then acknowledge whatever it demands
  local err seq gen
  err="$WORK/ack.err"
  FM_STATE_OVERRIDE="$STATE" "$DRAIN" >/dev/null 2>"$err" || return 1
  seq=$(sed -n 's/^WAKE_ACK_REQUIRED:.*--ack-through \([0-9][0-9]*\) --recovery-generation [A-Za-z0-9._-][A-Za-z0-9._-]*$/\1/p' "$err")
  gen=$(sed -n 's/^WAKE_ACK_REQUIRED:.*--ack-through [0-9][0-9]* --recovery-generation \([A-Za-z0-9._-][A-Za-z0-9._-]*\)$/\1/p' "$err")
  [ -n "$seq" ] && [ -n "$gen" ] || return 1
  FM_STATE_OVERRIDE="$STATE" "$DRAIN" --ack-through "$seq" --recovery-generation "$gen"
}

watch_bg() {  # <stdout-file>
  PATH="$FAKEBIN:$PATH" FM_STATE_OVERRIDE="$STATE" \
    FM_CREW_STATE_BIN="$FAKEBIN/fm-crew-state.sh" \
    FM_POLL=1 FM_SIGNAL_GRACE=1 FM_CHECK_INTERVAL=999999 FM_HEARTBEAT=999999 \
    "$WATCH" > "$1" 2>"$1.err" &
}

wait_for_exit() {  # <pid> <ticks>
  local pid=$1 limit=$2 i=0
  while [ "$i" -lt "$limit" ]; do
    kill -0 "$pid" 2>/dev/null || return 0
    sleep 0.1; i=$((i + 1))
  done
  kill "$pid" 2>/dev/null; wait "$pid" 2>/dev/null; return 1
}

wait_poll_cycle() {  # <pid>; 0 = alive through a whole poll cycle
  local pid=$1 limit=300 beat first now i=0
  beat="$STATE/.last-watcher-beat"; rm -f "$beat"; first=""
  while [ "$i" -lt "$limit" ]; do
    kill -0 "$pid" 2>/dev/null || return 1
    first=$(stat -c %Y "$beat" 2>/dev/null); [ -n "$first" ] && break
    sleep 0.1; i=$((i + 1))
  done
  while [ "$i" -lt "$limit" ]; do
    kill -0 "$pid" 2>/dev/null || return 1
    now=$(stat -c %Y "$beat" 2>/dev/null)
    [ -n "$now" ] && [ "$now" != "$first" ] && return 0
    sleep 0.1; i=$((i + 1))
  done
  return 1
}

reap() { kill "$1" 2>/dev/null || true; wait "$1" 2>/dev/null || true; }

send() {  # <resolve-key> <answer>
  # FM_GATE_REFUSE_BYPASS=1 is the harness seam fm-gate-refuse-lib.sh documents
  # for driving real fleet commands from a test context; without it fm-send
  # refuses because this validation runs inside a no-mistakes gate worktree.
  env -u NO_MISTAKES_GATE PATH="$FAKEBIN:$PATH" FM_GATE_REFUSE_BYPASS=1 \
    FM_ROOT_OVERRIDE="$HOMEDIR" FM_HOME="$HOMEDIR" \
    FM_SEND_LOG="$WORK/send.log" FM_SEND_SETTLE=0 \
    "$SEND" t1 --resolve-key "$1" "$2"
}

# ---------------------------------------------------------------------------
printf 'checkout under test: %s (%s)\n' "$ROOT" "$(git -C "$ROOT" rev-parse --short HEAD 2>/dev/null || echo '?')"
printf 'throwaway FM_HOME:   %s\n' "$HOMEDIR"

say "setup: a ship worker opens two captain decisions on task t1"
{
  printf 'window=sess:fm-t1\n'
  printf 'kind=ship\n'
} > "$STATE/t1.meta"
{
  printf 'needs-decision [key=budget]: approve the $400 spend?\n'
  printf 'needs-decision [key=vendor]: vendor A or vendor B?\n'
} > "$STATUS"
cat "$STATUS"

say "scenario 1: the worker's two decisions wake the supervisor once"
watch_bg "$WORK/watch1.out"; W1=$!
if wait_for_exit "$W1" 150; then
  if grep -qF "signal: $STATUS" "$WORK/watch1.out"; then
    ok "watcher exited and woke the supervisor: $(cat "$WORK/watch1.out")"
  else
    bad "watcher exited without naming the status signal: $(cat "$WORK/watch1.out")"
  fi
else
  bad "the worker's decisions never woke the supervisor"
fi

say "scenario 2: the supervisor sees both decisions and acknowledges the wake"
drain > "$WORK/drain1.out" 2>"$WORK/drain1.err" || true
cat "$WORK/drain1.out"
grep -qF '[key=budget]' "$WORK/drain1.out" && grep -qF '[key=vendor]' "$WORK/drain1.out" \
  && ok "both open decisions presented to the captain" \
  || bad "the drain did not present both open decisions"
ack_cycle && ok "wake acknowledged; queue is $( [ -s "$STATE/.wake-queue" ] && echo 'NOT empty' || echo empty)" \
  || bad "could not acknowledge the wake"

say "scenario 3: the supervisor answers BOTH decisions with two --resolve-key sends"
send budget "approved, go ahead" && ok "first answer sent (key=budget)" || bad "first --resolve-key send failed"
send vendor "go with vendor B"  && ok "second answer sent (key=vendor)" || bad "second --resolve-key send failed"
printf -- '--- status log after both answers ---\n'; cat "$STATUS"

say "scenario 4 (the fix): neither of this home's own answers wakes the supervisor again"
: > "$WORK/watch2.out"
watch_bg "$WORK/watch2.out"; W2=$!
if wait_poll_cycle "$W2"; then
  if [ -s "$WORK/watch2.out" ]; then
    bad "the home's own answers re-woke the supervisor: $(cat "$WORK/watch2.out")"
  elif [ -s "$STATE/.wake-queue" ]; then
    bad "the home's own answers queued a durable wake: $(cat "$STATE/.wake-queue")"
  else
    ok "watcher stayed asleep through a full poll cycle: no wake reason, empty wake queue"
  fi
else
  bad "the watcher EXITED on this home's own --resolve-key answers: $(cat "$WORK/watch2.out")"
fi

say "scenario 5 (adversarial): a later worker line on the same task still wakes"
printf 'blocked: need staging credentials to continue\n' >> "$STATUS"
if wait_for_exit "$W2" 150; then
  grep -qF "signal: $STATUS" "$WORK/watch2.out" \
    && ok "the worker's blocked: line woke the supervisor: $(cat "$WORK/watch2.out")" \
    || bad "the worker's blocked: line did not name the status signal: $(cat "$WORK/watch2.out")"
else
  bad "a worker line after two owned answers was SWALLOWED"
fi
reap "$W2" 2>/dev/null || true

say "scenario 6 (guard): the answers are still shown to the captain, not hidden"
drain > "$WORK/drain2.out" 2>"$WORK/drain2.err" || true
cat "$WORK/drain2.out"
if grep -qF 'resolved [key=budget]: answered: approved, go ahead' "$WORK/drain2.out" \
  && grep -qF 'resolved [key=vendor]: answered: go with vendor B' "$WORK/drain2.out"; then
  ok "both owned closes still printed on the captain-facing surface"
else
  bad "the ledger hid this home's own closes from the captain-facing surface"
fi
grep -qF 'blocked: need staging credentials' "$WORK/drain2.out" \
  && ok "the worker's blocker is presented too" || bad "the worker's blocker was not presented"

say "scenario 7 (adversarial): a real fm-teardown.sh leaves no orphaned ledger state"
# Stand up the fixtures the real teardown executable needs (a project clone with
# an origin, a task worktree with nothing unlanded, and the backend/PR probes it
# shells out to), then run bin/fm-teardown.sh itself on a task whose status log
# carries this home's own --resolve-key closes, i.e. a live ledger.
TD="$WORK/td"; TDSTATE="$TD/state"; TDBIN="$TD/fakebin"
mkdir -p "$TDSTATE" "$TD/config" "$TD/data" "$TDBIN"
for stub in treehouse tmux; do printf '#!/usr/bin/env bash\nexit 0\n' > "$TDBIN/$stub"; done
cat > "$TDBIN/gh-axi" <<'SH'
#!/usr/bin/env bash
case "${1:-} ${2:-}" in
  "pr list") printf '%s\n' "count: 0 (showing first 0)" "pull_requests[]: []"; exit 0 ;;
  "pr view") echo "error: pull request not found" >&2; exit 1 ;;
esac
exit 0
SH
cp "$TDBIN/gh-axi" "$TDBIN/gh"
printf '#!/usr/bin/env bash\nexit 0\n' > "$TDBIN/no-mistakes"
chmod +x "$TDBIN"/*
git init -q --bare "$TD/origin.git"
git -C "$TD/origin.git" symbolic-ref HEAD refs/heads/main
git clone -q "$TD/origin.git" "$TD/_seed" 2>/dev/null
git -C "$TD/_seed" -c user.email=t@t -c user.name=t commit -q --allow-empty -m baseline
git -C "$TD/_seed" push -q origin main
rm -rf "$TD/_seed"
git clone -q "$TD/origin.git" "$TD/project"
git -C "$TD/project" remote set-head origin main 2>/dev/null || true
git -C "$TD/project" worktree add -q -b fm/task-x1 "$TD/wt" main
touch "$TDSTATE/.last-watcher-beat"
{
  printf 'window=firstmate:fm-task-x1\n'
  printf 'endpoint_task_id=task-x1\n'
  printf 'worktree=%s\n' "$TD/wt"
  printf 'project=%s\n' "$TD/project"
  printf 'kind=ship\nmode=local-only\nspawn_gen=live-4907\n'
} > "$TDSTATE/task-x1.meta"
printf 'needs-decision [key=budget]: approve the spend?\n' > "$TDSTATE/task-x1.status"
env -u NO_MISTAKES_GATE PATH="$TDBIN:$PATH" FM_GATE_REFUSE_BYPASS=1 \
  FM_ROOT_OVERRIDE="$ROOT" FM_STATE_OVERRIDE="$TDSTATE" FM_HOME="$TD" \
  FM_SEND_LOG="$WORK/send2.log" FM_SEND_SETTLE=0 \
  "$SEND" task-x1 --resolve-key budget "approved" >/dev/null 2>&1 || true
if [ -s "$TDSTATE/.task-x1.home-appends" ]; then
  ok "precondition: the answered task carries a live home-appends ledger"
  ls -a "$TDSTATE" | grep -F home-appends
else
  bad "precondition: no ledger was written, so teardown has nothing to retire"
fi
mkdir -p "$TDSTATE/.task-x1.home-appends.lock"
printf '%s\n' 2147483646 > "$TDSTATE/.task-x1.home-appends.lock/pid"
env -u NO_MISTAKES_GATE FM_GATE_REFUSE_BYPASS=1 FM_ROOT_OVERRIDE="$ROOT" \
  FM_STATE_OVERRIDE="$TDSTATE" FM_DATA_OVERRIDE="$TD/data" FM_CONFIG_OVERRIDE="$TD/config" \
  PATH="$TDBIN:$PATH" "$ROOT/bin/fm-teardown.sh" task-x1 > "$WORK/td.out" 2> "$WORK/td.err"
TDRC=$?
printf 'fm-teardown.sh exit=%s\n' "$TDRC"
tail -n 5 "$WORK/td.out" "$WORK/td.err"
if [ "$TDRC" -ne 0 ]; then
  bad "the real teardown refused, so the ledger-retirement path was not exercised"
elif ls -a "$TDSTATE" | grep -qF 'home-appends'; then
  bad "teardown left ledger state behind: $(ls -a "$TDSTATE" | grep -F home-appends)"
else
  ok "real fm-teardown.sh removed the per-task ledger and its stale lock"
  printf 'remaining task-x1 state: %s\n' "$(ls -a "$TDSTATE" | grep -F task-x1 || echo none)"
fi

printf '\n===========================\n'
if [ "$FAILED" -eq 0 ]; then printf 'RESULT: all live scenarios passed\n'; else printf 'RESULT: at least one live scenario FAILED\n'; fi
printf 'artifacts under %s\n' "$WORK"
exit "$FAILED"
Evidence: 12 concurrent real --resolve-key answers: ledger stays well formed, every close presented

Source: 12 concurrent real --resolve-key answers: ledger stays well formed, every close presented

=== every send either closed its key exactly once or refused loudly
12 closed, 0 refused, out of 12 concurrent answers
PASS no close was silently lost or duplicated
PASS no torn or interleaved line in the status log

=== the ledger is well formed: ascending, non-overlapping, byte-accurate ranges
v1
ident=strong:33:7288885:2026-09-20 04:28:36.700591186 +0000
498 960
PASS ranges are ascending and non-overlapping
PASS no range runs past the log's 960 bytes

PASS every landed close is presented on the captain-facing surface
PASS exactly the 0 refused answers are still listed as open decisions
Evidence: The new regression test fails on base 1b1b6e0 and passes on target

Source: The new regression test fails on base 1b1b6e0 and passes on target

# The new regression test, run against BASE 1b1b6e0 (change reverted, test copied in) $ tests/fm-send-resolve-key.test.sh ok - fm-send --resolve-key: the answer send itself closes the open decision ok - fm-send --resolve-key: the close never re-wakes its own home, later lines still do not ok - the first --resolve-key answer was left to re-wake this home

# The new regression test, run against BASE 1b1b6e0 (change reverted, test copied in)
$ tests/fm-send-resolve-key.test.sh
ok - fm-send --resolve-key: the answer send itself closes the open decision
ok - fm-send --resolve-key: the close never re-wakes its own home, later lines still do
not ok - the first --resolve-key answer was left to re-wake this home

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

⏭️ **Rebase** - skipped

Step was skipped.

⚠️ **Review** - 1 info
  • ℹ️ bin/fm-classify-lib.sh:1395 - status_retire_presentation_task removes the new ledger lock with fm_lock_remove_path &#34;$home_appends_lock&#34; 2&gt;/dev/null || true, which discards the result, while the sibling ledger removal on line 1393 propagates failure into rc. Concrete sequence: a process is SIGKILLed while holding a directory-shaped lock and an unexpected file remains inside .&lt;id&gt;.home-appends.lock/; fm_lock_clean_known_files removes only the known names, rmdir then fails, the failure is swallowed, and status_retire_presentation_task returns 0 while AGENTS.md:145 states both the ledger and its lock are "removed by teardown". The residue is one empty-ish directory in state/ with no owner; the next fm_lock_try_acquire on that path still reclaims it through ordinary stale-owner recovery, so nothing is incorrect at runtime. Noted only because the round-3 remedy was specifically about exhaustive retirement of this pair; the || true is a deliberate match for the pre-existing fm_lock_remove_path call sites (bin/fm-wake-lib.sh:936, bin/fm-afk-start.sh:167), so treating it as an accepted tradeoff is reasonable.
✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 7 of 9 scenarios driven live against the product
Scenario Result Live Evidence
An in-flight watcher classification lands after the supervisor's --resolve-key answer; the supervisor must not be woken by its own close ✅ pass live live-inflight-classify-race.sh run against both commits (race-before-after.txt): base exits signal: .../t1.status, target stays asleep. The driver first asserts the race actually reproduced (class…
Two separate --resolve-key answers on one task wake the supervisor zero extra times ✅ pass live live-resolve-key-wake.sh scenario 4: after two real fm-send --resolve-key answers, the real fm-watch.sh survives a full poll cycle with no wake reason printed and an empty durable .wake-queue (live-…
Adversarial: a worker-authored line after the owned answers still wakes the supervisor ✅ pass live live-resolve-key-wake.sh scenario 5 and live-inflight-classify-race.sh step 8: appending blocked: need staging credentials makes the real watcher exit with signal: .../t1.status
Guard: the turn-ended historical annotation still presents this home's own close (recorded decision review-r2-4) ✅ pass live live-inflight-classify-race.sh step 7 on target: with the wake suppressed, the turn-end marker wakes the watcher and the drain prints `wake annotation: ... t1.status: resolved [key=budget]: answered…
Guard: UNREAD STATUS and the signal annotation still print both owned closes ✅ pass live live-resolve-key-wake.sh scenario 6: the drain prints both resolved [key=budget] and resolved [key=vendor] annotations alongside the worker's blocker
Adversarial: a real fm-teardown.sh leaves no orphaned home-appends ledger or lock ✅ pass live live-resolve-key-wake.sh scenario 7: a real project clone, worktree and task are torn down with bin/fm-teardown.sh task-x1 (exit 0) after a real --resolve-key answer wrote the ledger and a stale `…
Adversarial: simultaneous --resolve-key answers on one task keep the ledger well formed and lose no close ✅ pass live live-concurrent-answers.sh with 12 concurrent real fm-send processes, 5 consecutive target runs: single merged ascending range inside the log's byte bounds, no torn line, every landed close presente…
The change's regression test reproduces the reported failure: fails before the fix, passes after ⏸️ untested no The prior payload recorded this scenario as live=false: it rests only on running the unit test file tests/fm-send-resolve-key.test.sh (copied onto base 1b1b6e0 and on target), which is a test harnes…
Full tests/fm-watch-triage.test.sh regression pass ⏸️ untested no The suite is sleep-heavy and was starved by a parallel run of the same file in another no-mistakes worktree on this host, so it could not finish inside this step (98 of ~126 cases had passed with zero…
  • FM_LIVE_ROOT=&lt;target&gt; live-inflight-classify-race.sh and FM_LIVE_ROOT=&lt;base 1b1b6e0 export&gt; live-inflight-classify-race.sh - the live before/after race through the real fm-watch.sh, fm-send.sh and fm-wake-drain.sh
  • FM_LIVE_ROOT=&lt;target&gt; live-resolve-key-wake.sh - the live supervisor loop: worker decisions wake once, two real --resolve-key answers, no re-wake, worker blocker still wakes, real bin/fm-teardown.sh task-x1 retires the ledger
  • FM_LIVE_ROOT=&lt;target&gt; FM_CONC_N=12 live-concurrent-answers.sh (5 consecutive runs) plus 6 base runs for comparison - concurrent real fm-send --resolve-key processes against one task
  • tests/fm-send-resolve-key.test.sh on target (23 pass) and the same file copied onto base 1b1b6e0, where test_separate_resolve_key_answers_do_not_rewake fails
  • tests/fm-wake-queue.test.sh - includes test_separate_self_announced_answers_after_fold_are_owned, test_unreadable_status_is_not_owned, test_folded_worker_resolved_is_not_owned_lag, test_owned_growth_still_annotates_turn_ended
  • tests/fm-wake-drain-unread-status.test.sh - includes test_self_announced_pending_reply_close_still_surfaces and the retired-task-id ledger/lock removal case
  • tests/fm-watch-triage.test.sh - the change's cases (test_folded_worker_decision_without_home_append_still_wakes, test_separate_self_announced_answers_after_fold_wake_once, the self-announced-close cases) all passed early; the rest of the file was still running at 98 passes / 0 failures when I finished
⚠️ **Document** - 1 info
  • ℹ️ AGENTS.md:145 - Judgment call, left as-is: the clause "presentation is unaffected, so both the signal annotation and UNREAD STATUS still print those lines" in the AGENTS.md state inventory restates the same contract that docs/architecture.md:96 states ("it never removes a line from presentation, so both that annotation and the UNREAD STATUS section still print this home's own bookkeeping closes"). Both copies are currently true, so neither is stale. I judged this legitimate rather than duplication needing consolidation: docs/documentation-audiences.md classifies AGENTS.md as agent-runtime (an operating contract a Firstmate agent reads directly) and architecture.md as maintainer-architecture, and the surrounding AGENTS.md inventory entries follow the same pattern of appending one operative clause to each file's description. Flagging it only so the divergence risk is visible: if a future change alters whether the ledger touches presentation, both lines must move together, and architecture.md:96 is the owner.
✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

Self-announced bookkeeping appends now record their exact byte ranges.
Later drains and signal scans skip those ranges, so two distinct
--resolve-key answers after an OPEN DECISIONS fold do not each wake the
supervisor. Worker-authored lines outside that ledger still signal.
Keep the home-appends ledger so this home's own --resolve-key answers do not
each wake the supervisor, and keep fold-lag from substituting for watcher
classification. Take main's multi-line self-announced append so one answer can
close several keys in one write.

Conflicts resolved in bin/fm-wake-lib.sh, docs/architecture.md, and
tests/fm-watch-triage.test.sh.
@tiago-peixoto
tiago-peixoto deleted the fm/fm4907-sync branch September 20, 2026 04:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant